Conversation
IconId::Store points table.rs at assets/icons/store.a8, which only existed on the app-store branch, so every build of this branch failed to read it. Add the mask and its SVG source here so the table stands on its own.
text::line returns the drawn width, so the bare match evaluated to i32 where the function body expects (), failing the capsule build.
nonos-data/marketplace/index.bin has no make rule, so naming it as a hard prerequisite failed every build on a checkout without it (CI: No rule to make target). Wrapping it in $(wildcard) keeps the rebuild on a newer catalogue where it exists and drops the prerequisite where it does not.
The branch's manifest predated the switch to in-process Ed25519 and dropped the dependency while verify/crypto.rs imports it, so the capsule failed with an unresolved import. Restore main's manifest and add only the app_skeleton dependency boot_index.rs needs.
Twenty-five syscalls were published as caps = ["valid_token"] while the cap table demands a hardware, dev-root or time capability. Twenty-two take any one of Admin or a hardware capability, now published with caps_any; the dev-root and time calls need one capability, published with caps.
The cap table gates MDRO with MDRQ and MDRC on can_enrol_dev_root. The old syscall caps check cannot resolve that predicate and wants valid_token, so this fails it until abi/caps-check-fail-closed lands; the fixed check passes.
Every lane was pinned to -accel hvf -cpu host, which only macOS has, so no Linux host could boot an image. KVM when /dev/kvm opens read-write, hvf on macOS, TCG otherwise, the rule the boot matrix already uses; the display and audio backends follow the host too.
The trailer's magic picked the verifier for every root, so a Pedersen trailer was checked against the vendor root too. A local root's leaf is a commitment to a secret this kernel holds; the vendor root's is not, so there the trailer may no longer choose the weaker proof.
MkLocalSign let a LocalSign holder prove any capability it held, including LocalSign itself, so one signer could hand out the right to sign. A local proof now names nothing beyond AMBIENT_CAPS. crypto_proofs checks a minted trailer and its refusals against the kernel's own verifier files.
An Alpine package index is signed that way, and the capsule offered SHA-1 only without the prefix, so the index could not be checked at all. The rsa crate rebuilds the whole padded block and compares it. Nothing here signs, so offering SHA-1 verification mints nothing new with it.
resolve.rs names super::root, which the crate never mounted, so it did not build and nothing noticed because no workflow ran it. It mounts root.rs, follows the rename of absolute to visible, adds the /linux confinement and Phdr::file_range tests, and joins the proof-crate matrix.
A signature over a package covers the compressed bytes of one member, so a verifier needs to know where each one starts and ends. members() refuses a file with any byte that belongs to no member, where gunzip ends the stream and ignores the rest.
The index's .SIGN.RSA entry is checked against Alpine's x86_64 keys by the crypto service, the package's control member against the index's C: SHA-1, and its data against the control member's datahash. Only a Verified value reaches the store, so unauthenticated bytes are refused, not kept unvouched.
Without the operator-key rotation: NOX_OPERATOR_V1 keeps main's key, now read from .keys/marketplace_operator_ed25519.pub. The capsule embedded an index nothing built, so tools/nonos-market-index writes one, signed and verified when the operator seed is present and empty otherwise. nonos-mk moves to 78dae45 for the zk_trailer_hash field; the icon table is 49 long.
A release naming x86_64-linux counted as having its attestation, so any release could claim the exemption for itself. It now needs the linux. namespace too, which is where the store routes it: to the installer that authenticates the bytes before the machine mints their proof.
Conflicts were both sides adding: the install and app-store capsules sit together in Cargo.toml, mk and userspace, and init/mod.rs names the install queue once. The init loop and install queue are taken as the PR wrote them; the commit after this replaces its wake path.
The scheduler takes every ready process's priority lock from the timer interrupt. wake.rs and boost_init_for_drain took init's with interrupts on from syscall context, so a tick inside either spun forever on one CPU. The install queue now raises through the guarded setter the window queue uses.
MkAppInstall took a package name and the store's own readiness flag. It now takes a listing and release; init asks the market for readiness and the release's package hash, and the installer refuses bytes of any other BLAKE3. The index keeps each record's D: and p: lines, so dependencies come too.
…each CryptoMachineKey derives the machine key for any label a Crypto holder names, so a key the kernel keeps for itself needs a label no syscall can ask for. Kernel labels start with a zero byte, and the syscall refuses any label that does.
The local signing identity was random each boot, so consent was too. It is now derived from the machine key, and first-boot setup, which alone holds EnrolDevRoot, grants the local root as a named step and keeps a token only this machine can make; later boots restore it. The desktop profile now includes setup and the market, and builds every capsule it embeds.
MkAppLaunch queues a run for init, which spawns the personality to start the program the installer recorded outside /linux. Whether it may start is the exec gate's answer. The store drops its console-code enrolment, which it never held the capability for, and gains an o key to open.
Every boot resolved validator.nymtech.net through net.dns, so the first query of a session went out in the clear naming the service in use. The bootstrap set is pinned by address in this capsule's attested image; the name is kept for the certificate check and never resolved.
The installer connected to dl-cdn.alpinelinux.org, which the socket service resolved in the clear. It now takes a mirror by address, Alpine's CDN by default or NONOS_ALPINE_MIRROR at build, sends Alpine's name as the Host line, and refuses a host that is not a literal. The signatures still decide.
The generator hashed packages from v3.21 while the installer downloads from v3.20, so every Linux listing pinned bytes the installer never sees and the kernel's package-hash check would refuse every install.
The spawn installs an empty token, and nothing on the console said so. The line reads the bits back from the process table rather than printing the value it meant to install, so it is evidence and not a restatement.
The four answered at once whatever their timeout, so a program waiting on descriptors spun, and poll(NULL, 0, ms), the idiom for a short sleep, did not sleep. select also narrowed the caller's sets even when nothing was ready, and copied a whole 1024-bit set whatever nfds was. poll counted an entry with a negative descriptor, which Linux ignores and programs use to switch an entry off, as closed and so ready. Each now waits with the other parked calls until something it watches is ready or its timeout passes: poll's int of milliseconds, negative for none; ppoll's and pselect6's timespec and select's timeval, null for none, with negative or out-of-range fields refused with EINVAL. A timeout of zero only looks. select writes its sets back only when something is ready, copies only the longs nfds covers, always empties the exception set, and empties all three when its time runs out, as Linux does. A negative poll descriptor is never ready. A wait watching a socket or a timer is looked at again every 10 ms, now for poll and select as for epoll.
Two parts added to cwait. Two threads block reading one empty pipe and both must be answered once bytes arrive. poll, ppoll and select each wait out a 100 ms timeout with nothing ready, poll(NULL, 0, 50) sleeps, and a select with no timeout is woken by another thread's write, with its set narrowed to the ready descriptor. Both pass on host Linux.
…aces A closed descriptor stayed in every epoll interest list. Linux drops it, so a program that closes without EPOLL_CTL_DEL and gets the same number back from its next open adds it again; here that add was refused with EEXIST, and the old entry reported the new file under the old token. dup2 onto an open descriptor overwrote it without closing it, so a file's buffered bytes were lost and a socket's handle kept, and a target number past the table grew the table to reach it. close now drops the descriptor from the guest's interest lists. dup2 closes an open target first, as Linux does, and refuses a number past the table with EBADF.
A twelfth part: a watched pipe end is closed without EPOLL_CTL_DEL, a new pipe takes its number, and adding it again succeeds with no stale event. dup2 to a number past the table is EBADF. Passes on host Linux.
A timerfd fired once and forgot its interval, so a periodic timer stopped after its first read, and a read always answered one. timerfd_create ignored its clock and its TFD_NONBLOCK and TFD_CLOEXEC flags, timerfd_settime ignored TFD_TIMER_ABSTIME and never wrote the old setting, timerfd_gettime was not served, and a read of a timer that had not fired answered EAGAIN even to a blocking descriptor. The expiry lived in the descriptor, so dup and fork gave a timer that never fired. A timer is now an object the family keeps, lent with the pipes and eventfd counters, so every descriptor onto it sees one timer. It is made on a clock Linux names (anything else is EINVAL) with its flags; set relative or absolute on its own clock, one-shot or periodic, with the old setting written when asked; read for how many times it has fired since the last read, stepping a periodic timer past them; read blocking until it fires unless the descriptor is non-blocking; and reported readable once it has fired. A wait watching a timer is woken when it fires rather than on a 10 ms look.
A thirteenth part: an unarmed non-blocking timer reads EAGAIN, a 50 ms periodic timer counts its firings over 180 ms and reports its interval, a one-shot read blocks until it fires, and an absolute time 100 ms ahead wakes an epoll wait. Passes on host Linux.
gopreempt spins a goroutine with no call in it beside main, on one P. Go moves such a goroutine off the CPU only by sending SIGURG to its thread while it runs, so main waking from a 20 ms sleep and a garbage collection, which stops the world, both depend on that signal landing. On host Linux main runs again after 25 ms and the collection ends by 45 ms, with the tgkill deliveries visible under strace.
ioctl answered ENOTTY to every request, so a program asking how many bytes a pipe or file held, setting a descriptor non-blocking the BSD way, or marking it close-on-exec without fcntl was refused. FIONREAD answers what a pipe's read end holds, or what is left of a file past its offset; FIONBIO sets or clears O_NONBLOCK from the int it is given; FIOCLEX and FIONCLEX set and clear close-on-exec. A socket's FIONREAD, and every device request, is still ENOTTY: there is no terminal or device behind a descriptor here, which is also how isatty says no.
A duplicated or inherited epoll descriptor started with an empty interest list, so a child that waited on the epoll its parent set up before fork waited on nothing. The list is now copied with the descriptor. Linux shares one list between them; the copy holds what was registered at the dup or fork, and a change made afterwards on one side is not seen on the other.
A fourteenth part: FIONREAD counts three bytes in a pipe, FIONBIO makes its read end answer EAGAIN once drained, FIOCLEX sets close-on-exec, and a forked child finds a ready entry in the epoll list its parent built. Passes on host Linux.
sched_getscheduler, sched_setscheduler, sched_getparam, sched_setparam, sched_get_priority_max and _min and sched_setaffinity were not served, nor epoll_create and epoll_pwait2, so a program that sets a thread's policy or pins it, or a libc that reaches for the older or newer epoll form, died there. They are answered as Linux answers an unprivileged process on the one CPU the guest is shown: every thread is SCHED_OTHER at priority zero; asking for SCHED_OTHER, BATCH or IDLE at zero is accepted and changes nothing, since scheduling is the kernel's; asking for FIFO or RR is EINVAL outside the priority range 1 to 99 and EPERM inside it, the range checked first as Linux checks it, since no guest holds the privilege Linux asks for; the priority ranges are Linux's. An affinity mask that includes the one CPU is accepted and one that leaves it out is EINVAL. A pid argument is mapped from the guest's namespace like kill's, and one outside the family is ESRCH. epoll_create checks its size is positive and makes a list; epoll_pwait2 waits like epoll_pwait with a timespec timeout.
A fifteenth part: the policy is SCHED_OTHER, SCHED_FIFO at priority zero is EINVAL, SCHED_OTHER is accepted, the real-time range is 1 to 99, CPU 0 can be pinned, epoll_create refuses a size of zero, and epoll_pwait2 waits its 50 ms timespec. SCHED_FIFO at priority one and a CPU-1-only mask depend on privilege and CPU count, so they are printed, not checked. Passes on host Linux.
gettid and getpid answer with the family's own numbers, and kill and tkill map theirs back, but tgkill's were passed through as they came. A thread signalling itself with tgkill(getpid(), gettid(), sig), which is how Go's runtime preempts a goroutine and how glibc's pthread_kill and raise reach a thread, named numbers no kernel thread has and was answered ESRCH. Both of tgkill's pid arguments are now mapped, and one outside the family is ESRCH, as for kill.
A sixteenth part: a caught SIGUSR1 sent with tgkill(getpid(), gettid()) runs its handler, and a thread the family does not have is ESRCH. Passes on host Linux.
cwait stopped at its first failing part, so a run with more than one change removed showed only the first. Every part now runs, each failure is printed, and the last line counts them. A hang still stops it at the part it hangs in.
A signal reached a guest thread only as the answer to a call it had made, so a thread running its own code never received one. Go preempts a goroutine that spins without a call by sending SIGURG to its thread, and a guest is shown one CPU, so such a goroutine held the only P for good: no other goroutine ran and a garbage collection never stopped the world. SIGALRM, SIGINT and any other signal for a busy thread waited the same way. MkForeignInterrupt lets a supervisor mark one of its guest threads. The timer trampoline, after a tick that interrupted user mode, parks a marked thread with its whole register file as a frame numbered NR_INTERRUPTED and wakes the supervisor; the thread sleeps as a parked call does. A signal answer rewrites the trampoline's frame to enter the handler, keeping the thread's FPU state for its return as a delivered call does; any other answer lets the thread run on exactly where it was. A thread already parked in a call is not marked, since that call's answer can carry the signal. Only the supervisor recorded for the thread may mark it; the mark is dropped with the thread; the handler's context is checked as for any signal answer. The trampoline's frame is its 160 bytes and no more: it is read and written as those 20 words, never as the whole SavedUser, whose TLS words would lie past the top of the kernel stack. A tick with nothing marked costs one load. The libc gains mk_foreign_interrupt and FOREIGN_NR_INTERRUPTED.
A caught signal raised for another thread waited in the queue until that thread next made a call, so a thread running its own code never saw it. Raising one now also asks the kernel to stop the target at its next tick. The stopped thread arrives as a FOREIGN_NR_INTERRUPTED frame and is delivered what is pending, with its handler returning to the thread's own rax since no call is being answered; with nothing pending it runs on where it was. A thread parked in a call, the caller included, still gets the signal with that call's answer.
cpreempt spins a thread on a flag only its SIGUSR1 handler sets, sends it SIGUSR1 with pthread_kill, and joins it. The thread makes no call while it spins, so the join returns only if the handler runs inside the spin. It passes on host Linux, pinned to one CPU as well.
The frame a handler is entered on wrote the 18 saved registers at ucontext + 48 and the blocked mask right after them. Linux, musl, glibc and Go all read uc_mcontext at ucontext + 40 and uc_sigmask at + 296 (measured with offsetof on musl and glibc). Every register a handler read from its context was therefore the one before it: Go's preemption handler reads rip and rewrites rip and rsp, so it would have read the saved rsp as the pc and resumed into garbage. The C handlers so far never read their context, and the frame tests only read it back the way it was written, so nothing showed it. The registers now start at + 40 and the mask sits at + 296. rt_sigreturn reads the same constant, so a frame still round-trips, and a new test pins r8, rsp, rip and the mask to the offsets musl and glibc use. With the old offset that test fails and the four round-trip tests pass.
… handlers on it sigaltstack read back an empty stack and ignored the one a program set, so every handler was entered on whatever stack the thread was using. Go installs every handler with SA_ONSTACK and gives each thread its own signal stack; its handler checks that it runs on that stack or on g0's, and a frame on a goroutine's stack sends it into a path that waits for a spare M a program without cgo never has. That is where gopreempt stopped: one SIGURG delivered to tid 50, then no further line in 420 s. Each thread's alternate stack is now kept, and sigaltstack answers as Linux's do_sigaltstack: the old setting is what was there before the call, SS_ONSTACK while the thread runs on it and SS_DISABLE when none is set; changing it from on it is EPERM, a stack under MINSIGSTKSZ (2048) is ENOMEM, and a mode other than 0, SS_ONSTACK or SS_DISABLE is EINVAL. SS_AUTODISARM is named unserved and answered EINVAL, as a Linux before 4.7 answers it. A handler with SA_ONSTACK is entered at the top of the alternate stack when the thread is not already on it, and below the interrupted rsp when it is; a frame that would run off the bottom is not written. uc_stack carries the stack and its flags for the interrupted rsp. A fork gives the child the forking thread's stack, a thread's exit drops its own, and execve clears them all.
A seventeenth part sets a 16 KiB alternate stack through the raw call, so the answers are the kernel's and not musl's own checks, and raises a SIGUSR2 caught with SA_ONSTACK. It passes only if: none set reads back SS_DISABLE; 1024 bytes is ENOMEM and flags 4 is EINVAL; the stack set reads back with flags 0; the handler's own local is on that stack while the saved rsp is not; inside the handler the stack reads SS_ONSTACK and changing it is EPERM; uc_stack names it with flags 0; and SS_DISABLE turns it off again. On the container's Linux it prints "cwait altstack ok: ... flags inside 1" and "cwait PASS: 17 parts".
MkForeignInterrupt marks a running thread so that the next tick in user mode stops it and hands it to its supervisor. When the thread makes a call before that tick, the call's answer is where the supervisor delivers the signal, but the mark stayed, and a later tick stopped the thread once more with nothing to deliver. The instrumented gopreempt run showed it: after "signal 23 to tid 50" on a call's return, the next tick took tid 50 again and the supervisor answered it with a plain resume. A thread's mark is now dropped as it traps into a call. The check costs one atomic load while nothing is marked.
clone returned the new thread's tid in the guest's pid namespace, but wrote the kernel's pid for it to the CLONE_PARENT_SETTID and CLONE_CHILD_SETTID words. musl keeps the parent's copy as the thread's own tid and hands it to tkill, so pthread_kill asked for a tid the guest does not have and got ESRCH. An instrumented cpreempt showed it, "[C] dbg kill rc 3", then a join that waited for a signal never sent. gettid and Go's tgkill were right, since both use the translated number. The family now writes those words once the reply has been put in the guest's terms, with the same number clone returns. clone itself keeps recording CLONE_CHILD_CLEARTID, which is keyed by the kernel's pid.
An eighteenth part starts a thread that yields until a SIGUSR1 handler sets its flag, then sends pthread_kill to it and joins. It passes only if pthread_kill answers 0 and the handler ran. The thread gives up after 3 s, and sooner when the kill failed, so a lost signal fails the part and does not hang it. On the container's Linux it prints "cwait pthread-kill ok ... handled 1" and "cwait PASS: 18 parts".
This was referenced Sep 29, 2026
eKisNonos
added a commit
that referenced
this pull request
Sep 30, 2026
#582 gained nine commits on 30 September. Four are already here as the same commits. The other five are earlier forms of 07d9c92, d7b2019, f5fe356, 2b5c726 and 5b40bbc, which this branch carries in the split form linux/next-waits gave them. The tree here already holds all of it, so the merge keeps it as it is, and #582 and this branch now merge into main in either order.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Targets
main. It also carries the 209 commits of #567 (linux/store-install), which come before it and hold the Linux personality these commits change, so it merges after #567 and shows those commits until #567 lands. What it adds is thirty-six commits.Before this, a Linux-guest test image started no guest at all. Once it did, no guest could run a second thread: Go's threads called address zero, musl's
pthread_createfailed outright, and a thread's exit returned to it. A guest that did get threads reset the machine when it exited, left its threads running, and hung if one of them faulted. And nothing could wait:eventfd2was not served, so every Go program with a timer died at its first one, whileepoll_waitand a futex with a timeout returned at once or never.What Linux guests can do now
time.Sleep, a timer woken early through Go's eventfd, and a pipe read through the poller to end of file.epoll_wait,pollandselecttimeouts, non-blocking pipes, several readers or a writer waiting on one pipe, edge-triggered epoll, and timerfd one-shot, periodic and absolute.epoll_create,epoll_pwait2andtgkillare served as Linux serves them to an unprivileged process on one CPU.What changed
The boot guest starts
bf9d9dbf3: the Linux-guest test image (NONOS_LINUX_GUESTS=1) buildsmicrokernel-desktop-guiwithnonos-stark-attest, without first-boot setup. Under the setup profile every app, the personality included, waits for setup to exit, and setup waits for keys, so the unattended image never started its guest: 0[APP-LINUX]or[LINUX]lines in 180 s. Every other build ofnonos-mk-desktop-gui-prodis unchanged.Threads start, run and end
3fc1ac19d: a clone child starts on a copy of the calling thread's registers, with rax 0, the new stack, and the parent's thread pointer unless a TLS is named.MkForeignThreadtakes the calling thread as a fifth argument; it must be supervised by the caller and in the same thread group as the target, so no guest receives another's registers. Zero keeps the fresh start. Go's child calls r12 and musl's calls r9, so both called address 0 before. A TLS is taken only whenCLONE_SETTLSis set.abi/syscalls.tomlrecords the new argument.addec49af: clone writes the new tid whereCLONE_PARENT_SETTIDandCLONE_CHILD_SETTIDask. musl's thread-list lock stores its owner's tid, and a thread that never learned its tid locked the list as 0, which the lock reads as free.1e6b8a673:mprotectcommits a PROT_NONE reservation it is asked to open. musl reserves every pthread stack with PROT_NONE and opens the usable part withmprotect;MkPeerProtectreprotects only pages that exist, sopthread_createfailed. A backed region is reprotected as before; a reservation is backed with the asked protection, keeping any page the guest already touched, and recorded as backed so fork copies it. A span outside every region is refused with ENOMEM, as on Linux.10947aa96: a thread'sexitends the thread and never returns to it. The word it named throughCLONE_CHILD_CLEARTIDorset_tid_addressis zeroed and one waiter woken, as Linux does. musl names its thread-list lock there andpthread_joinwaits for that lock, so every join hung before. Go'sexitThreadfalls into INT3 whenexitreturns, which the fault report below would have turned into the end of the process.A process ends whole
5ab0a0912: a thread group's page tables stay until its last member leaves the process table. Release freed them when the leader was finalized, under a Go thread still running on them, and the machine reset: CR3 was the guest's, FS held the Go thread's TLS, CR2 was the IDT entry for vector 8. The tables now pass to a process still running on them, whose capability token is rebound to the ASID it now owns.874f3ecfc:MkKilladmits the caller the foreign registry names as the target's supervisor. A thread's parent is its leader, not the supervisor, so every thread kill at a guest's exit returned EPERM and the threads ran on. A refused kill now prints[LINUX] kill refused: pid <n> outlives its process, errno <e>.2311afc39: when a guest thread ends on a signal, and only when it ends itself so a supervisor's own kill cannot loop back, the kernel leaves a one-way notice for its supervisor, collected on its nextMkForeignWaitas a frame numberedFOREIGN_NR_DIED. The kernel reports; what a dead thread means is the personality's policy.959b57f98: the personality ends the process on that notice, as Linux ends a thread group on an unhandled fault, and reports the signal as 128+signo.Fork
ffc1b542b: a fork copies the calling thread's parked frame, and the kernel carries that thread's own thread pointer to the child. Before, every fork took the leader's registers and the personality's singlefs_base, which is only the last thread to set one.Waiting, as Linux waits
69f3b48d4: a pipe has Linux's ends. The family notes, each time it lends its pipes to the guest being answered, which ends are still open anywhere in the family. From that note: a read end is readable while it holds bytes and hung up once no write end is left; a write end is writable while there is room and in error once no read end is left. An empty pipe with no write end reads as end of file, and a write no reader can take is refused with EPIPE. poll, epoll and select report hang-up and error unasked. Before, every pipe was readable and writable to all three.3e48edbaf: a descriptor keeps O_NONBLOCK. F_SETFL sets it, F_GETFL reports it with the access mode, pipe2 takes it, dup and fork carry it, and a non-blocking read of an empty pipe answers EAGAIN. Before, F_SETFL was dropped, which is how Go's runtime prepares every pipe for its poller.bfb07892f: eventfd2 and eventfd are served. The counter is kept with the pipes as the family's, so dup, fork and every thread reach one. EFD_SEMAPHORE, EFD_NONBLOCK and EFD_CLOEXEC as Linux; all ones and a short buffer EINVAL. Go's runtime makes one at its first timer and threw when it could not.06a266ac5: epoll_wait and epoll_pwait wait for their timeout (an int of ms; negative waits until ready). A blocking eventfd read or write and a write to a full pipe wait too. A waiting call is left parked in its trap and tried again after every answer and at its deadline, so a write from another thread ends the wait at once. A wait on a socket or timer is looked at every 10 ms. A pipe write of up to PIPE_BUF goes in whole or waits. maxevents of zero or less is EINVAL.32d2a1a06: EPOLLET entries are reported as readiness rises and re-armed by an EAGAIN; EPOLLONESHOT reports once. epoll_ctl refuses as Linux: EBADF, EPERM for a regular file or directory, EINVAL, EEXIST, ENOENT. Go registers every pipe and socket with EPOLLET and EPOLLOUT, so a level-triggered report of an always-writable end would answer its poller at once, every time.aa0087d75: a futex wait ends with ETIMEDOUT at its timeout. FUTEX_WAIT_BITSET (absolute, MONOTONIC or REALTIME), FUTEX_WAKE_BITSET, FUTEX_REQUEUE and FUTEX_CMP_REQUEUE are served. The value is compared as 32 bits, so musl's sign-extended -1 matches. musl's timed waits and condition variables, and Go's sysmon, use these.b1ead15c5: two proof guests, gopoll and cwait (below).More waiting, timers, descriptors and threads
b716a34af: any number of threads can wait to read one pipe. A read parked in one slot per process, so a second reader of an empty pipe took the slot and the first was never answered. Pipe reads now wait with the other parked calls; the family settles them after it reaps, so a writer that left with its process reads as end of file at once. A zero-byte read answers zero.439ed66c6:poll,ppoll,selectandpselect6wait for their timeout (poll's int of ms, ppoll's and pselect6's timespec, select's timeval; null or negative waits until ready; bad fields EINVAL).poll(NULL, 0, ms)sleeps. select writes its sets back only when something is ready, copies only the longs nfds covers, always empties the exception set, and empties all three when its time runs out. A negative poll descriptor is never ready.85b3d7724: a closed descriptor leaves every epoll interest list, as on Linux, so its number can be added again; before, that add was EEXIST and the old entry reported the new file under the old token.dup2closes an open target first (a file's buffered bytes were lost, a socket's handle kept) and refuses a number past the table.b9efc7f96: timerfd has Linux's semantics and is kept with the family like the eventfd counter, so dup and fork share it. A periodic timer keeps firing and a read returns how many times it fired; the clock is Linux's set (else EINVAL); TFD_NONBLOCK, TFD_CLOEXEC and TFD_TIMER_ABSTIME are honoured; the old setting is written when asked;timerfd_gettimeis served; a blocking read waits for the firing; a wait watching a timer is woken when it fires.262c06280: FIONREAD (a pipe's bytes, a file's remainder), FIONBIO, FIOCLEX and FIONCLEX. Any other request is ENOTTY, as for a descriptor with no terminal or device behind it.8f778075b: a dup'd or inherited epoll descriptor carries the interest list it had; before, it started empty.ff778c704:sched_getscheduler,sched_setscheduler,sched_getparam,sched_setparam,sched_get_priority_maxand_min,sched_setaffinity,epoll_create,epoll_pwait2. Every thread is SCHED_OTHER at zero; SCHED_FIFO and SCHED_RR are EINVAL outside 1 to 99 and EPERM inside it, the range checked first as Linux checks it; a mask without the one CPU is EINVAL; pid arguments are mapped like kill's.9039b893d: tgkill's thread group and thread are mapped into the family's pids. gettid and getpid answer with the family's numbers, sotgkill(getpid(), gettid(), sig), which Go's runtime and glibc'spthread_killuse, was ESRCH.84e10f538,7c3967f18,468f5aa97,2f0db3b25,a045bec06,66b55faf8,afd96ef58,24cafe103: the guest gopreempt, and seven more cwait parts, one per change above; cwait now runs every part and names each one that fails.Cleanups
289272374: deletecall/clone.rs, a copy left behind when clone moved tocall/spawn/; no module included it.ce395f7ff: the[PF] demand fillline is printed only after a fill, not before the handler refuses the null page or the kernel half.Evidence
q35 under TCG, one vCPU. Each guest is the boot guest of its own store.
Threads, faults and exit (the first thirteen commits)
[GO] hello PASS[GO] conc PASS: 8 goroutines summed 3199960000[C] cthreads PASS: 8 pthreads joined, summed 3199960000[LINUX] guest thread 51 ended on a signal; ending the process, then[LINUX] guest exitedthreadfault's worker runs on a stack mapped read-write outright through
clone(), so it does not depend on the pthread path. ItsSURVIVEDline, printed only if main outlives the faulting thread, never appears.Waiting (the next seven commits)
Both guests pass on host Linux, which is what they are measured against. On one build carrying all twenty commits:
[GO] poll PASS: 3 parts, 3454ms in all: slept 69 ms, the early timer ended in 18 ms, the pipe read to end of file in 339 ms[C] cwait PASS: 9 parts in 951 ms: cond_timedwait ETIMEDOUT after 102 ms, a blocking eventfd read woken after 147 ms, epoll_wait timed out after 100 ms and woken after 153 ms, POLLHUP on a closed pipe, a full-pipe write waited 148 ms, edge-triggered counts as LinuxThe same build ran the first four guests again: gohello 283 calls, goconc 253, cthreads 87, threadfault 7 with its death line; 0 unserved in each.
More waiting, timers, descriptors and threads (the next sixteen commits)
On one build carrying all thirty-six commits:
[C] cwait PASS: 16 parts in 2313 ms; SCHED_FIFO at priority 1 errno 1 and a CPU-1-only mask errno 22; a 50 ms periodic timer fired 3 times in 180 ms, a one-shot read waited 103 ms, an absolute timer woke epoll after 105 ms[GO] poll PASS: 3 parts, 3437ms in allThe same build ran the regression set: gohello 282 calls, goconc 248, cthreads 87, threadfault 7 with its death line; 0 unserved in each.
Each fix was also run with its change removed, on the same guest:
MkKillkill refused ... errno 1lines; guest threads fault afterguest exitedSURVIVEDprinted, no death lineSURVIVEDmprotectcommit (the pthread build of threadfault, before the fix)[C] threadfault FAIL: no workerpthread_createsucceeds (cthreads)a 10ms timer set under a 3s one took 2823ms, FAILeventfd2served[LINUX] unserved nr=290, thenfatal error: runtime: eventfd failedFAIL: edge-triggered counts (222, 1): the second look counts 2, not 1FAIL: end of file after the write end closed (0, 0): poll sees nothingFAIL: 6 parts failed, 10 passed: poll and ppoll returned after 2 ms, the re-add was EEXIST, the periodic timer counted 1, FIONREAD failed,[LINUX] unserved nr=145, tgkill ESRCHFAIL: epoll list through fork (13, 256): the child finds nothing ready and exits 1Each of these ran on its own build with only that change removed, except the six that fail different cwait parts, which were removed together. gopoll passes with the hang-up removed, as its writer closes before its reader waits, and with O_NONBLOCK dropped, as Go's pipe reads then park instead of polling.
The store settled in 12.8 to 84.2 s across the 22 guest boots for the first thirteen commits, in 38.0 to 104.8 s across the 16 for the next seven, and in 17.2 to 117.7 s across the 10 for the sixteen after; the personality waits up to 300 s.
Checks
check_stubs,check_allows,check_dark_features: 0 new sites.check_syscall_abi: 111 published syscalls reach a handler.check_unreachable: 1 new site,has_children, whichlinux/store-installreports too; this branch has 1257 sites against its 1258.microkernel-core) and the personality were checked at every one, with 0 errors. The next twenty-three change only the personality and the guests; the personality was checked at each with 0 errors and 0 warnings, and each guest passes on host Linux.Not done here
mprotecton a backed region does not update the recorded protection, so a fork after it gives the child the protection the region was mapped with.exit(notexit_group) still ends the whole process; on Linux the others would run on.killto a child process is queued in the sender, not the child, so the child's handler never runs; an uncaught one ends the child as it should.